Back

Neural Networks

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Neural Networks's content profile, based on 35 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Dendritic Wave Recurrent Neural Networks

Kubo, Y.

2026-07-09 neuroscience 10.64898/2026.07.03.736415 medRxiv
Top 0.1%
7.9%
Show abstract

Wave recurrent neural networks (wRNNs) are biologically inspired recurrent architectures that use traveling-wave dynamics to support sequence learning and memory. However, their input-to-hidden pathway remains relatively simple compared with biological neurons, where dendrites perform nonlinear input integration. In this study, we introduce the Dendritic Wave Recurrent Neural Network (DWRNN), which augments the input pathway of the wRNN with nonlinear basal dendritic branches while preserving the original recurrent wave dynamics. We evaluate DW-RNN on a simple copy task, sequential MNIST (sMNIST), permuted sequential MNIST (psMNIST), and noisy sequential CIFAR-10 (nsCIFAR-10). On the copy task, DW-RNN shows learning behavior comparable to the standard wRNN, suggesting that dendritic input integration does not disrupt the recurrent wave-based memory mechanism. On the three sequential image-classification benchmarks, DW-RNN outperforms the standard wRNN, improving accuracy from 97.27 {+/-} 0.15% to 97.82 {+/-} 0.12% on sMNIST, from 96.74 {+/-} 0.17% to 96.92 {+/-} 0.10% on psMNIST, and from 54.30 {+/-} 0.79% to 55.65 {+/-} 0.55% on nsCIFAR-10. In addition to improving mean accuracy, DW-RNN exhibits lower across-seed variability on all three classification benchmarks, suggesting that dendritic input integration may improve the stability of wRNN training. Hidden-activity visualizations further show that DW-RNN preserves the characteristic traveling-wave patterns of the original wRNN. These results suggest that dendritic computation and traveling-wave recurrent dynamics provide complementary mechanisms for biologically inspired sequence learning.

2
Structural Composition Enables Very Fast Learning

Riveland, R.; Pouget, A.; Latham, P.

2026-07-15 neuroscience 10.64898/2026.07.14.738142 medRxiv
Top 0.3%
1.9%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWThere is a gap between neuroscientific theories of learning and the speed of learning observed in many experiments. Since the Cognitive Revolution of the 1950s, compositionality has played a central role in efforts to bridge this gap. Roughly, a compositional system is one where distinct modules are combined according to a set of rules in order to accomplish complex tasks. Recently, significant progress has been made in understanding the emergence of modules in both biological and artificial neural systems. How, and under what conditions, the rules of module recombination are represented in these systems remains an open question. Here we present a neural model that can leverage these rules to dramatically speed up learning. We first show that when faced with multiple tasks which share subcomponents, models learn a low-dimensional representation that captures how subcomponents are reused across the task set. These low-dimensional spaces encode the structure that governs how modules should be recombined. Restricting learning to these subspaces greatly reduces the amount of experience needed to acquire a novel task, even when learning from reinforcement on single trials. In some cases, we can leverage the geometric regularities of these representations to reduce learning to a form of hypothesis testing over a small set of discrete points. Finally, we use this theory to model both behavioral and neural data from non-human primates performing a compositional task, and show that key features in this data are consistent with a model in which exploration during learning is restricted to these low-dimensional spaces. Overall, this work shows that the advantages of modularity in neural systems can be greatly improved upon when models represent the structure of module reuse. Both these features working in tandem lead to learning on timescales similar to biological intelligences, and hence provide a model for how such fast, adaptable behavior can emerge from systems of neurons.

3
Context-Aware Evidence-Gated Plasticity for Multi-Goal Learning in Spiking Neural Networks

Neymotin, S. A.; Hazan, H.; Unal, G.; Earl, C.; Anwar, H.; Franaszczuk, P.; Boothe, D.

2026-06-30 neuroscience 10.64898/2026.06.25.734613 medRxiv
Top 0.3%
1.7%
Show abstract

Background / Introduction: Biologically inspired spiking neural networks can model adaptive behavior, but learning multiple goals is difficult because synaptic updates for different targets can interfere. We tested whether multi-timescale plasticity and context-specific credit assignment could improve continual multi-goal learning in a spiking navigation system inspired by entorhinal-hippocampal circuitry. Methods: We developed a closed-loop spiking model containing grid-like, place-like, target-related, association, and motor-output populations. An agent navigated in a two-dimensional environment with randomized starting locations and learned through reward-modulated spike-timing dependent plasticity (STDP/RL) and a novel evidence-gated plasticity (EGP) framework. EGP accumulates candidate synaptic modifications, evaluates them using reward evidence, and consolidates only changes that improve performance. A target-context variant maintained separate proposal stores and reward evaluation for each target. Results: STDP/RL learned and retained a single-target navigation policy, but multi-target training produced substantial interference, including attraction to incorrect targets after learning. Across 10 connectivity seeds, target-context EGP achieved higher late-stage reward than global EGP, improved weakest-target performance, and increased the fraction of targets achieving positive reward. In a longer continual-learning simulation, reward increased for all targets, TEST-phase performance increasingly exceeded TRAIN-phase performance, and proposal magnitudes grew over learning. Dwell-time confusion analyses showed that target-context EGP reduced wrong-target attraction and improved target selectivity relative to multi-target STDP/RL. Conclusions: These results demonstrate that spiking navigation circuits can learn goal-directed behavior using local plasticity, but robust multi-goal learning benefits from context-specific evidence-based consolidation. Target-context EGP provides a biologically motivated mechanism for reducing interference during continual reinforcement learning in spiking neural networks.

4
A neural network model of free recall learns multiple memory strategies

Li, M.; Jensen, K. T.; Zhang, Q.; Lu, Q.; Mattar, M. G.

2026-07-06 neuroscience 10.1101/2025.09.25.678592 medRxiv
Top 0.3%
1.7%
Show abstract

Humans exhibit structured patterns of memory recall, including a tendency to recall more recent information and to recall events in the same order they were experienced. Classic computational models explain these patterns by positing that memories incorporate the ongoing ''temporal context'', formed by smoothly integrating the stimulus history. However, it is unclear whether a single mechanism can account for the full repertoire of human memory strategies, as the optimal approach may be task-dependent. For example, human memory experts widely apply the ''memory palace'' strategy, which is empirically better but not captured by temporal context models. Here we show that neural networks optimized for free recall develop diverse retrieval strategies, with only some of them resembling temporal context models.The best-performing models discovered a stimulus-invariant index code that emphasizes the studied position of each list item, instead of its temporal context. This creates a stable scaffold for forward recall akin to the memory palace technique. This index code was more likely to emerge when networks were i) encouraged to recall all studied items rather than prioritizing a few items, and ii) prevented from relying on recency, resonating with human data. Our findings demonstrate that human-like recall patterns can arise from multiple distinct computational mechanisms, and that sequential retrieval using item index is an optimal strategy that explains expert-level recall performance.

5
Spatiotemporal transformation of neural data reveals representations of erroneous behaviors

Sihn, D.; Kim, S.-P.

2026-07-04 neuroscience 10.64898/2026.07.04.736476 medRxiv
Top 0.4%
1.4%
Show abstract

Abnormal states such as erroneous behaviors are generally difficult to represent from neural data. However, such states are also known to have specific spatiotemporal features, indicating a feasibility of developing a method to focus on them. If a method can highlight these spatiotemporal features, it may effectively represent such abnormal states, helping evaluate abnormal brain functions. In the present study, we proposed the hierarchy of supported modules (HSM) to highlight spatiotemporal features that can represent abnormal states. HSM spatiotemporally transforms multidimensional neural time-series based on their spatiotemporal context. We evaluated HSM through decoding and similarity analyses using multiple publicly available datasets. In the HSM results, decoding accuracies were higher for erroneous behaviors than for normal behaviors, and similarities were lower between erroneous behaviors and normal behaviors than between normal behaviors, demonstrating the ability of HSM to capture the spatiotemporal features of erroneous behaviors. Surprisingly, many parts of these results were also present even before HSM learning, showing the virtue of HSM as a simple-to-use method. The proposed HSM method may help elucidate the mechanisms underlying erroneous behaviors.

6
Preserved geometry during representational drift enables stable perception and memory

Zaid, H.; Schaffer, E. S.

2026-06-28 neuroscience 10.64898/2026.06.25.734656 medRxiv
Top 0.4%
1.3%
Show abstract

In many brain regions, the stimulus tuning of neurons is stable on a timescale of hours but not on a timescale of weeks, a phenomenon often called representational drift. This would seem to imply that these brain regions cannot be used for stable recognition of sensory stimuli or the retrieval of associative memories learned several weeks prior. However, decoding approaches have demonstrated that in some cases, stable decoding of drifting representations is possible. In principle, adaptive decoding provides a plausible resolution to the paradox of how the brain operates with drifting representations, but we lack a deep understanding of what the requirements are for stable decoding to be possible. Here, we offer a general mathematical framework that explains when and why stable decoding from a drifting representation can be achieved. First, we demonstrate that both feedforward and recurrent networks preserve the geometry of their inputs when the network is sufficiently large, meaning that representational drift must also preserve geometry in these networks. Second, we demonstrate that drifting representations that have stable geometry are decodable with adaptive decoders. Therefore, not only the existence of preserved geometry in the presence of representational drift but also the ability to decode from drifting representations simply requires the population of neurons exhibiting representational drift to be large. This theoretical framework not only suggests that preserved geometry should be a general feature of drifting representations, it also explains the conditions under which empirical efforts to measure stable geometry will be successful.

7
Learning complex temporal dependencies via local synaptic plasticity

Ng-Kee-Kwong, J.; Tang, M.; Akam, T.; Bogacz, R.

2026-07-10 neuroscience 10.64898/2026.07.09.737423 medRxiv
Top 0.5%
1.1%
Show abstract

The ability to extract and exploit temporal structure across diverse tasks is central to human cognition. Neuroscientists have typically relied on recurrent neural networks (RNNs) trained with backpropagation through time (BPTT) when modelling neural and behavioural processes such as decision-making and motor control. However, this algorithm has limited biological plausibility, hence the computational principles underlying efficient learning of temporal dependencies remain unresolved. Here, we investigate temporal predictive coding (tPC), a recently proposed framework that extends predictive coding to the temporal domain while preserving local Hebbian update rules. We analyse and extend tPC to establish its relationship with several influential computational models of learning in RNNs, including BPTT, reservoir computing, and eligibility propagation (e-prop). We first demonstrate a functional equivalence between tPC and tBPTT1, a variant of BPTT in which gradients are propagated only one time step into the past. We then show that tPC can leverage reservoir dynamics to encode short-range temporal context, and simultaneously sculpt neural trajectories in state space to support downstream readout. We further demonstrate that hierarchical recurrent dynamics can facilitate learning of more complex temporal dependencies, while additionally conferring robustness to strong distractors. Finally, we show that tPC networks can be augmented with biologically inspired eligibility traces to solve temporally extended context-dependent tasks. Together, these results reveal that relatively simple recurrent networks governed by local plasticity can support temporal learning in more complex settings than previously appreciated.

8
Topological data analysis captures complex behavioral dynamics during naturalistic social interaction between domestic ferrets

Reiling, J.; Padilla-Coreano, N.; Patel, D.; Frohlich, F.; Zhang, M.

2026-07-07 neuroscience 10.64898/2026.07.01.735818 medRxiv
Top 0.5%
1.0%
Show abstract

Capturing naturalistic behavioral dynamics is essential for understanding social interaction in ecologically valid settings. Existing investigations of naturalistic social interaction rely on time-aggregated analysis methods better suited for task-based experiments, which lose the complex, moment-to-moment dynamics exhibited in naturalistic settings. The emerging field of topological data analysis (TDA) provides new tools to characterize fine-grained dynamics in time-series data that cannot be captured by time-averaged methods. The present work utilizes Temporal Mapper, a recently developed TDA specifically tailored to analyzing dynamical systems. Temporal Mapper characterizes complex temporal dynamics as transition networks, where nodes are stable states and edges are transitions between states. Originally designed for human neural time series analysis, here we demonstrate the utility of Temporal Mapper to capture rich animal postural dynamics during naturalistic social interaction. We utilized an existing dataset with 12 video recording sessions of two domestic ferrets (Mustela putorius furo) during naturalistic interaction and tracked the postures of animals during social interaction. Ferrets were chosen due to their strong social-cognitive skills and rich postural dynamics for investigating social behavior via posture estimation. Temporal Mapper was then used to represent the postural dynamics as transition networks for each recording session. Here, we found that posture states are significantly smaller and more widespread during active social interaction compared to non-social activities. Additionally, the number of sequential postural states before transitioning to new behaviors is more consistent during active social interaction than non-social activities. Together, our findings suggest that social activity has a broad range of unstable postural states arranged in consistent sequences. Our method, Temporal Mapper, allows for network structure analysis of complex naturalistic data, applicable for characterizing rich dynamics in different species, scales, and paradigms.

9
Quantitative assessment of mesoscale cellular order and organization in the mouse hippocampus

Hein, K. O. R.; Romero-Limon, H.; Moeckel, C.; Karasinsky, A.; Kayser, J.; Moellmert, S.; Zaccone, A.; Guck, J.; Toda, T.

2026-07-09 neuroscience 10.64898/2026.07.06.736467 medRxiv
Top 0.6%
0.9%
Show abstract

The hippocampus is characterized by a stereotypical macroscopic structure, where the nuclei are densely and heterogeneously packed among different subregions of the hippocampus. Despite the fact that tissue-specific cellular organization has been implicated in neural function, it has been technically challenging to quantitatively analyze mesoscopic cellular organization in the hippocampus due to its high cellular density. To overcome this technical hurdle, we developed Computational Biophysical Histomorphometry Software (CBHS), an automated image-analysis pipeline, aimed at quantifying nuclear shape and the order of the cellular ensemble in high-density areas. When applied to the subfields of hippocampus, we found that denser regions, most notably the dentate gyrus, were the most positionally, but least orientationally ordered. Nuclear shape exhibited a dependence on the local environment in a packing-dependent manner. This association was cell-type specific, with neurons, but not astrocytes displaying nuclear shape that varied with neighbour proximity, although astrocytes demonstrated greater intrinsic shape variance. The results reveal the presence of reproducible mesoscale cell packing order in hippocampal tissue, and are consistent with a nucleus-driven mechanical coupling between neighbouring cells. The present study provides a quantitative framework with which to understand mesoscopic tissue organization, thus enabling the formulation of testable hypotheses for future investigation.

10
Transitive reasoning as linear classification

Ferrera, V. P.; Lippl, S.; Kay, K.; Munoz, F.; Jin, Y.; Jensen, G.; Terrace, H.

2026-06-28 neuroscience 10.64898/2026.06.24.734346 medRxiv
Top 0.6%
0.9%
Show abstract

Transitive inference (TI) is the ability to reason about transitive relationships in an ordered set of items (e.g., if A>B and B>C, then A>C). TI is widely held to depend on a linear representation of the serial (rank) order of those items. By what computational mechanism is such an ordering constructed during learning, and how is it used to make choices that obey transitivity? Here we take a minimalist approach, applying least-squares estimation (LSE) to a serial learning task commonly used to test TI in humans and animals. In this formulation, LSE computes a linear classifier that maps task conditions onto behavioral outcomes. This algorithm makes no explicit assumptions about transitivity or serial order, yet it reproduces key empirical features of TI; namely, the ability to generalize beyond the training set, and a symbolic distance effect (SDE) in performance accuracy. Applying the classifier to individual items produces an internally ordered representation of rank from which both generalization and the SDE naturally emerge. The approach also yields a decision mechanism, in the form of a differencing operation, for selecting the correct item from any pair. These findings reframe TI as a linear classification problem, challenging conventional assumptions about the cognitive mechanisms required for transitive reasoning.

11
Charge-trap flash memory cells of the brain

Foster, P. P.; Chhikara, R. S.; Boriek, A. M.

2026-07-03 neuroscience 10.64898/2026.06.29.733154 medRxiv
Top 0.6%
0.8%
Show abstract

Despite extensive study of cellular mechanisms underlying long-term potentiation, no single specific protein or gene has been identified which encodes an individual unit of information, or memory bit. Indeed, the brain engram remains a knowledge gap. The theory of exclusion led us to cancel one-by-one several unrealistic biological options, suggesting that the explanation resides somewhere else. Superposition of up to concentric 300 myelin layers, spiraled, and highly compacted wrapping a single axon and each wrap could host hundreds to thousands of niches, as memory cells, collectively consisting of a massive array of cells. The disjointed 3D spatial superposition allows storage of charges, nodes not facing from a layer to next. The thickness of a single myelin layer ranges from 7.0 to 20 nm. The dimension scale is approximately the exact dimensions of the charge trap, the tunnel and dielectric also equipping current AI microchips. Stored charges are positive ions, with similar effect whether charges are negative or positive charges creating an electromagnetic field. To write data, following an action potential, this voltage applies to the control gates of the myelin layers producing an ionic charge injection. This causes charges to gain energy and tunnel through the myelin layer across Ranvier nodes, via quantum tunneling, and deep into the concentric myelin multilayers. This is creating an insulated trapping of K+ ions isolated from the system. In a long white matter tract bundle, the near-perfect isolation of millions of axons within compressed myelin wrap-ion channel K+/Na+ systems provides quantum coherence and precision of asynchronous firing property. The injected ionic charges (K+) become physically stuck in traps within the myelin layers. The K+ ions may not move freely, completely trapped after AP ceases. Mirroring a single-bit, single-level-cell, a trapped ionic charge (ions K+) may represent a 1, while an empty cell (absence of K+) represents a 0. The trial-and-error process, with a Bayesian inference which may have also been the core evolution of the learning human brain. Based on selected mathematical equations, we analyzed the general scheme on how deep learning may be embedded in the brain

12
Self-Supervised Behavioral Representations Across the Life Course: A Killifish Case Study

Chang, J.-C.; Komatsu, T. S.; Onami, S.

2026-06-29 animal behavior and cognition 10.64898/2026.06.23.733896 medRxiv
Top 0.6%
0.8%
Show abstract

Self-supervised foundation models of aging are increasingly built from longitudinal data (biobanks, electronic health records, wearables) that is inherently incomplete: no individual is followed across a whole lifetime, and how much of each life is captured varies widely. This raises two linked questions: is it worth modeling an individual's whole life course rather than its current state, and can such a model be built from brief, fragmentary records? No human cohort can settle them, because none offers a complete life to compare against. We turn to the African turquoise killifish (Nothobranchius furzeri), tracked from youth to natural death in publicly released recordings, as a controlled testbed: its complete lifespans provide the full-life reference that human data lacks. On these data we build LifeMAE, a two-stage selfsupervised model: a day encoder that summarizes each day of behavior, then a life-course encoder over the trajectory of those daily summaries. We find that the day encoder alone is already strong: from a single day of behavior it predicts chronological age, separates long- from short-lived individuals (coarsely), and flags nearness to death. Adding the life-course encoder improves on none of the three; each is matched by trivially aggregating the day-level predictions (a smoother for age, an early-life average for lifespan). Near-term mortality seems the exception, where the whole-life model looks far better (AUROC 0.81 to 0.91), but the gain is not behavioral: it reflects where each day falls within the observation window (a cue supplied by the model's encoding of time), and a single-day model given that cue closes the gap at any observation length. For these traits, an individual's place in its life course is legible from a single day: the trajectory stage is unnecessary, and the record it needs is as short as one day, the finest grain our day-level setup resolves. For characterizing a cohort, this favors observing many individuals briefly over tracking a few for long. The result joins a growing body of work in which deep and foundation models, fairly benchmarked, fail to beat deliberately simple baselines. We add a concrete mechanism for the over-optimism: a model's encoding of time can leak the very quantity it predicts, which backwardlooking evaluation mistakes for learned biology, so only evaluation fixed to the moment of prediction is trustworthy.

13
Learning the wiring rules of a mammalian cortical column

Richter, O.; Schneidman, E.

2026-07-10 neuroscience 10.64898/2026.07.09.737432 medRxiv
Top 0.6%
0.8%
Show abstract

Characterization of neural circuits' architecture typically relies on measurable neuronal features such as morphology, molecular identity, and spatial location. While generative models leveraging these properties have proven accurate, they remain constrained by available measurements and our assumptions regarding the prospective features. Here, we present an alternative approach using representational learning and use it to model the circuitry of a column of the mouse primary visual cortex. Our framework learns jointly low-dimensional embeddings of neurons in an abstract feature space alongside wiring rules that predict synaptic connectivity. These embedding-based models accurately predict individual synapses, connectivity degrees, and network motif statistics -- outperforming standard generative models that depend on detailed cell-type classifications -- using only a handful of embedding dimensions and wiring rules. Crucially, the learned representations prove interpretable, recapitulating cortical depth, cell type, and dendritic morphology. The resulting wiring blueprint is both simple and biologically meaningful, suggesting that cortical connectivity follows surprisingly parsimonious logic. This framework offers a general and exportable tool for learning minimal generative models of connectomes.

14
PolliCrop: A high-throughput computer vision pipeline for pollinator monitoring in agroecosystems

Chabert, S.; Bernigaud-Samatan, J.; Blackman, B. K.; Blanchet, N.; Catrice, O.; Donnadieu, C.; Gani, M.; Grousset, R.; Husband, S.; Tueux, G.; Erler, S.; Langlade, N. B.

2026-07-13 animal behavior and cognition 10.64898/2026.07.08.737348 medRxiv
Top 0.7%
0.6%
Show abstract

Flower-visiting insect populations are declining since the 1990s, especially because of the decrease of floral resources in agricultural settings. Mass flowering crops can help increase resource availability, and plant breeding can be directed towards selecting varieties attracting more flower-visiting insects. This requires the implementation of an automated high-throughput phenotyping tool for assessing the attractiveness of plant genotypes to flower-visiting insects. In this study, (i) we present a procedure to take standardized images of sunflower heads with camera traps continuously at day and night in the field; (ii) we trained two versions of a deep learning model, named PolliCrop, to automatically detect and identify three classes of the main insects visiting sunflower on these images (non-Bombus bees, bumble bees, lepidopterans); (iii) we assessed and validated the ability of PolliCrop to correctly predict the true visitation frequencies of the insect classes on three sunflower genotypes; (iv) we presented two statistical approaches to compare the insect visitation frequencies between plant genotypes, one including weather variables, and the other one without. One PolliCrop version yielded satisfying performance to correctly detect the three insect classes. In particular, it correctly predicted the insect visitation frequencies on two sunflower genotypes in a range of {+/-}10%. The other PolliCrop version can be useful in certain contexts of images and objectives. PolliCrop can be extended in the future to other crop species by training PolliCrop on new images captured in these crops. The field experimental design to set up for comparing the attractiveness between genotypes is also discussed.

15
GCBM-DCT-HV-Bio-NL-Grow-CHG-CSM-RHEC: A Unified Geometric, Biological, Causal, and Regenerative Framework for Mechanism-Aware Tissue and Connectome Modeling

Xu, T.; Hu, Z.; Sun, X.; Jin, L.; Xiong, M.

2026-06-29 bioinformatics 10.64898/2026.06.24.734320 medRxiv
Top 0.7%
0.6%
Show abstract

Modern biological prediction problems increasingly require models that go beyond Euclidean feature regression and local graph smoothing. Tissue, cellular, and connectome systems are nonlinear, geometry-dependent, intervention-sensitive, history-dependent, and subject to regenerative or homeostatic constraints. We propose GCBM/DCT/HV/Bio/NL/Grow/CHG/CSM/RHEC, a unified model for mechanism-aware biological prediction. The model integrates geometric connectome dynamics, differentiable charted tissue geometry, Hamiltonian latent transport, nonlinear biological kinetics, nested latent memory, continual growth without overwriting, causal hypergraph structure, causal structure modeling, and regenerative homeostatic error correction. Unlike Euclidean baselines, which treat observations as flat vectors, and local graph baselines, which use neighborhood smoothing without mechanistic structure, the proposed model represents biological states (Trapnell 2015) as coupled geometric, dynamical, causal, and regenerative objects. We evaluate the model on four synthetic toy studies, Toy A, B,C, D, designed to reflect increasing biological complexity: local Euclidean structure, nonlinear mechano-chemical interaction, causal intervention response, and out-of-distribution regenerative shift. Compared with Euclidean and local graph baselines, the full model achieves the lowest mean squared error across all four toy studies. Relative to the Euclidean baseline, the full model reduces MSE by approximately 63.0%, 89.1%, 89.0%, and 90.9% on Toy A, Toy B, Toy C, and Toy D, respectively. These results support the value of integrating geometry, mechanism, causal structure, adaptive growth, and regenerative correction into a single predictive architecture (Figure 1).

16
Uncovering internal states with a robust shared-state multi-neuron GLM-HMM framework

Lawrence, A.; Yezerets, E.; Janak, P. H.; Charles, A.

2026-07-02 neuroscience 10.64898/2026.06.27.734988 medRxiv
Top 0.7%
0.6%
Show abstract

Neural systems exhibit multiple firing states that reflect an organism's internal state and modulate the relationship between external environmental stimuli and behavior. Several studies have inferred these latent states by supplementing the traditional hidden Markov Model (HMM) with generalized linear models (GLMs) with non-Poisson behavioral observations. However, understanding the relationship between internal brain states and behavior also requires modeling the neural activity. Nonetheless, fitting multi-neuron GLM-HMMs is non-trivial due to high sparsity, collinearity, and low trial counts in neuronal datasets. Therefore, we built a robust multi-neuron GLM-HMM framework that uncovers latent states from population activity while incorporating the influence of time-stamped task variables and spike histories. To obtain reliable model parameters, we employ a modified expectation-maximization procedure. Specifically, we show that incorporating neuron-adaptive penalization in the maximization step overcomes the covariate co-linearity issues typical of time-stamped events and sparse spiking, yielding stable estimates of Poisson GLM coefficients. Furthermore, we incorporate a trust-region algorithm to ensure stable M-step convergence in the presence of ill-conditioned Hessians that can lead to unstable Newton-Raphson updates. We further demonstrate the utility of leave-one-out cross-validation analysis for evaluating model performance on datasets with low trial counts and without breaking their temporal structure. We evaluate our framework on three electrophysiological datasets from primates and rodents as they perform a decision-making task, demonstrate stable model convergence, and discuss the behavioral relevance of the inferred states.

17
Data-driven oscillatory network modeling with condition-dependent coupling laws: Identifying directed neural interactions in working memory attention dynamics

Ohkawa, M.; Zhou, Y. J.; Haegens, S.; Jafarian, M.

2026-07-10 neuroscience 10.64898/2026.07.06.736523 medRxiv
Top 0.7%
0.6%
Show abstract

Learning new information in the presence of distracters and changing conditions requires the ability to adapt. In the brain, this adaptive capability has been linked to dynamic interactions between attention and working memory, which enable the selective filtering of irrelevant input while preserving behaviorally relevant information. Specific neural oscillations have been implicated in this process. Here, we introduce a phenomenological data-driven framework for oscillatory network modeling that learns condition-dependent coupling laws directly from neural recordings and enables inference of condition-dependent directed pathways. We apply our approach to magnetoen-cephalography (MEG) data collected while participants performed a working-memory task with and without distracters. Recall dynamics in the non-distracter condition are first modeled using a linear oscillatory network in which each region of interest is represented by two alpha-band harmonic oscillators. We use universal differential equations (UDE), an extension of neural differential equations, to capture distracter-induced changes in coupling laws. Symbolic regression is then used to interpret the modifications identified by UDE as nonlinear functions, and an additional method is proposed to identify the directed pathway from the newly emerging nonlinear terms in the dynamics of brain regions of interest. Despite inter-subject variability, working memory recall data from all four participants examined under distraction showed the emergence of a pathway from the dorsolateral prefrontal cortex (dlPFC) to the primary visual cortex (V1). This finding is consistent with the established role of the dlPFC in cognitive control and suggests that distracter processing recruits a directed interaction from prefrontal to visual regions. More broadly, our results illustrate that combining linear models whose parameters are learned from the data with universal differential equations augmented by interpretability methods enables the identification of condition-dependent coupling laws, their representation as interpretable mathematical functions, and the discovery of candidate directed pathways underlying adaptive changes in oscillatory networks without requiring strong prior assumptions about the underlying mechanisms.

18
A single dynamical property can account for the capacity to learn, from artificial networks to the mammalian brain.

Hengen, K. B.; Chopra, R.; Zhong, J.; Miller, E. S.; Bekele Tolossa, G.; Fosque, L. J.; Meza, J. A.; DeKorver, N. W.; Guerriero, R.; Ritter, N. J.; Lambo, M. E.; Bhaskaran-Nair, K.; Van Hooser, S. D.; Shew, W.

2026-07-10 neuroscience 10.64898/2026.07.09.737603 medRxiv
Top 0.8%
0.5%
Show abstract

Every brain must adapt to an unpredictable world, yet individuals differ in how readily they learn. Theoretical work suggests that learning is fastest when a system, whether biological or synthetic, is initialized in a state close to instability - i.e., near criticality - because critical dynamics are imbued with a diverse repertoire of patterns and multi-scale correlations. Here, we empirically estimate distance to criticality in the brain and show that it predicts the rate of adaptability underlying learning, neuronal tuning, and general intelligence. In mouse motor cortex, proximity to criticality forecasts learning rate of two future complex tasks: prey capture hunt and ladder crossing. In contrast, distance to criticality predicted neither an animal's naive ability nor its asymptotic skill - isolating the rate of learning itself. In visual cortex of young ferrets, proximity to criticality predicts how strongly experience reshapes neural tuning. In human frontal cortex, it correlates with general cognitive ability. A minimal recurrent network model reproduced these results and offers a mechanism: proximity to criticality defines the timescale over which a system can learn from its past experiences, directly setting the rate of learning. A single dynamical property can account for the capacity to learn, from artificial networks to the mammalian brain.

19
Interpretable compositional computation with recurrent neural networks

Pezon, L.; Van Meegen, A.

2026-06-29 neuroscience 10.64898/2026.06.23.733979 medRxiv
Top 0.8%
0.5%
Show abstract

Flexible cognition utilizes reusable components to enable rapid adaptation of behavior to different contexts or tasks. Analysis of artificial neural networks trained on multiple tasks suggested that this compositionality is supported by dynamical structures which are shared and re-used across tasks. However, the nature of these shared components, and how they can be used in a task-dependent manner, remained unclear. Here, we develop a theory of interpretable compositional computation based on shared dynamical structures in the low-dimensional latent space of low-rank recurrent neural networks. We show that these shared latent components are not immediately visible in the neural activity, and are thus compatible with task-dependent activity. We identify hallmarks of shared latent components both in the connectivity statistics and the neural representations. These hallmarks yield testable predictions for the networks response to specific perturbation experiments. Finally, we identify distinct loci where task-dependence can enter the computation, allowing us to characterize qualitatively different solutions to compositional tasks. In summary, our theory provides a mechanistic understanding and testable hallmarks of compositional computation via shared components in low-rank networks.

20
Modelling individual ampullary afferents in two species of gymnotiform fish using simulation-based inference

Mayer, S.; Benda, J.; Grewe, J.

2026-06-30 neuroscience 10.64898/2026.06.24.734418 medRxiv
Top 0.8%
0.5%
Show abstract

Ampullary electroreceptors are widespread across aquatic vertebrates. The purpose of sensing exogeneous electric fields is conserved across species but the implementations differ and the encoding mechanisms remain incompletely understood. We compared baseline and stimulus-driven response properties of ampullary electroreceptor afferents in the weakly electric fish Apteronotus leptorhynchus and Eigenmannia virescens. We find that their activity is well captured by an extended leaky integrate-and-fire model that generalizes across both species. The model shares similarities to a previous model of the tuberous electroreceptor afferents but further incorporates a low-pass pre-filtering and additional noise sources to reproduce the observed spectral response characteristics. The low-pass is essential to shape stimulus encoding in the high-frequency range. Accurate prediction of low-frequency stimulus encoding further requires two distinct noise sources: stimulus-independent white current noise and activity-dependent noise in the adaptation current, which is shaped by the adaptation time constant to yield effective pink noise dynamics. Using simulation-based inference, we trained a neural network to map model parameters to neuronal response features. This approach enables the generation of heterogeneous, biologically plausible model populations that may serve as a realistic input layer for studying neuronal processing on the next level. With this, we provide a unified and mechanistic model of ampullary electroreceptor encoding in these species and possibly beyond.